Back

Nature Computational Science

Springer Science and Business Media LLC

All preprints, ranked by how well they match Nature Computational Science's content profile, based on 55 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Integrative Inference of Spatially Resolved Cell Lineage Trees using LineageMap

Pan, X.; Chen, Y.; Zhang, X.

2026-01-22 developmental biology 10.64898/2026.01.19.700383 medRxiv
Top 0.1%
30.2%
Show abstract

Understanding the spatio-temporal processes of tissue growth, including how new cell types emerge and how cells form the tissue architecture, is a fundamental problem in biology. The emerging spatially resolved lineage tracing data, where three modalities, lineage barcodes, gene expression profiles, and spatial locations, are measured for each single cell, provides an unprecedented opportunity to understand these processes. Computational methods that take advantage of all three modalities to reconstruct cell lineage tree and ancestral cell states and locations are needed. We introduce LineageMap, a hybrid lineage inference algorithm that integrates the scalability of distance-based tree reconstruction methods with the flexibility of likelihood-based methods under a unified probabilistic framework. The input to LineageMap is spatially resolved lineage tracing data, where for each single cell, the gene expression, lineage barcode and spatial locations are available. LineageMap enables accurate, interpretable, and scalable inference of high-resolution lineage trees as well as locations of ancestral cells from the tri-modality single-cell data. Across simulated and experimental datasets, LineageMap consistently outperforms existing methods in the accuracy of reconstructed cell lineage trees, while revealing biologically coherent spatiotemporal trajectories. Our framework bridges molecular lineage tracing with spatial and transcriptomic information, advancing computational reconstruction of dynamic cellular ancestries in both time and space. LineageMap is available at: https://github.com/ZhangLabGT/LineageMap.

2
Virtual Epileptic Patient (VEP): Data-driven probabilistic personalized brain modeling in drug-resistant epilepsy

Wang, H. E.; Woodman, M.; Triebkorn, P.; Lemarechal, J.-D.; Jha, J.; Dollomaja, B.; Vattikonda, A. N.; Sip, V.; Medina Villalon, S.; Hashemi, M.; Guye, M.; Scholly, J.; Bartolomei, F.; Jirsa, V.

2022-01-21 neurology 10.1101/2022.01.19.22269404 medRxiv
Top 0.1%
26.5%
Show abstract

One-third of 50 million epilepsy patients worldwide suffer from drug resistant epilepsy and are candidates for surgery. Precise estimates of the epileptogenic zone networks (EZNs) are crucial for planning intervention strategies. Here, we present the Virtual Epileptic Patient (VEP), a multimodal probabilistic modeling framework for personalized end-to-end analysis of brain imaging data of drug resistant epilepsy patients. The VEP uses data-driven, personalized virtual brain models derived from patient-specific anatomical (such as T1-MRI, DW-MRI, and CT scan) and functional data (such as stereo-EEG). It employs Markov Chain Monte Carlo (MCMC) and optimization methods from Bayesian inference to estimate a patients EZN while considering robustness, convergence, sensor sensitivity, and identifiability diagnostics. We describe both high-resolution neural field simulations and a low-resolution neural mass model inversion. The VEP workflow was evaluated retrospectively with 53 epilepsy patients and is now being used in an ongoing clinical trial (EPINOV).

3
Spectral normative modeling of brain structure

Mansour L, S.; Di Biase, M. A.; Yan, H.; Xue, A.; Venketasubramanian, N.; Chong, E.; Alexander-Bloch, A.; Chen, C.; Zhou, J. H.; Yeo, B. T. T.; Zalesky, A.

2025-01-21 radiology and imaging 10.1101/2025.01.16.25320639 medRxiv
Top 0.1%
18.2%
Show abstract

Normative modeling in neuroscience aims to characterize interindividual variation in brain phenotypes and thus establish reference ranges, or brain charts, against which individual brains can be compared. Normative models are typically limited to coarse spatial scales due to computational constraints, limiting their spatial specificity. They additionally depend on fixed regions from fixed parcellation atlases, restricting their adaptability to alternative parcellation schemes. To overcome these key limitations, we propose spectral normative modeling (SNM), which leverages brain eigenmodes for efficient spatial reconstruction to generate normative ranges for arbitrary new regions of interest. Benchmarking against conventional counterparts, SNM achieves a 98.3% speedup in computing accurate normative ranges across spatial scales, from millimeters to the whole brain. We demonstrate its utility by elucidating high-resolution individual cortical atrophy patterns and characterizing the heterogeneous nature of neurodegeneration in Alzheimers disease. SNM lays the groundwork for a new generation of spatially precise brain charts, offering substantial potential to drive advances in individualized precision medicine.

4
Tau accumulation patterns in PSP constrain mechanisms and quantify cell-to-cell and cell-autonomous aggregation rates.

Huang, S.-H.; Quaegebeur, A.; Pansuwan, T.; Rittman, T.; Wang, R.; Knowles, T. P.; Rowe, J. B.; Klenerman, D.; Meisl, G.

2024-12-20 neurology 10.1101/2024.12.14.24318991 medRxiv
Top 0.1%
14.8%
Show abstract

Protein aggregates are a hallmark of neurodegenerative disease, yet the molecular processes that control their appearance are still poorly understood. In particular, it is unknown to what degree the development of aggregates in one cell is triggered by nearby aggregated cells, as opposed to cell-autonomous processes. Here we develop a cell-level computational model to test alternative hypotheses of disease progression from human data and demonstrate its applicability in the primary tauopathy Progressive Supranuclear Palsy. From brain slices stained for aggregated tau, we quantify the contribution of cell-to-cell and cell-autonomous processes to the proliferation of aggregates across different brain regions and disease stages. We find that the triggering of aggregation by nearby aggregated cells, over distances in the order of 100{micro}m, is the major driver of disease progression. Our computational model can then simulate interventions to evaluate potential therapeutic strategies in a virtual reconstruction of a human primary neurodegenerative tauopathy. HighlightsO_LIA minimal mathematical model can reproduce tissue-level aggregate accumulation patterns in silico at cellular resolution. C_LIO_LICell-to-cell interactions determine aggregate patterns in progressive supranuclear palsy (PSP). C_LIO_LICell-to-cell interactions are not limited to nearest neighbours, but act over a millimetre-scale. C_LIO_LIReducing cell-to-cell interactions or cell vulnerability, rather than targeting cell-autonomous processes, is a potential disease-modifying therapeutic strategy. C_LI

5
CeDNe: A multi-scale computational framework for modeling structure-function relationships in the C. elegans nervous system

Moza, S.; Zhang, Y.

2025-11-04 neuroscience 10.1101/2025.11.03.683805 medRxiv
Top 0.1%
12.8%
Show abstract

Understanding how neural circuits generate behavior requires integrating structural and functional data across scales. C. elegans with its complete connectome, genetically identifiable neurons, single-cell transcriptome, neuropeptide-receptor distribution, and an amenability to simultaneous measurement of brain-wide neural activity and behavior presents a unique opportunity for such a multiscale circuit analysis. However, the absence of a unifying framework to connect these diverse datasets limits our ability to connect network structure and attributes with function. Here we introduce CeDNe (C. elegans Dynamical Network), an open-source computational framework that integrates anatomical, molecular, and imaging datasets into a unified graph-based representation that enables multimodal data analysis by cross-referencing different omics layers in a single computational environment. Specifically, CeDNe provides modular tools for visualizing and analyzing network connectivity, motif distribution, and circuit paths. Further, it incorporates a computational framework that simulates neural dynamics and optimizes network models to bridge structural connectivity with neural activity. Thus, CeDNe establishes a scalable foundation for data-driven modeling of the nervous system. This open-source tool not only facilitates computational connectomics and multimodal analyses in C. elegans but also serves as a generalizable framework for investigating structure-function relationships in neural networks of other organisms.

6
Data-driven motif discovery in biological neural networks

Matelsky, J. K.; Robinette, M. R.; Wester, B.; Gray-Roncal, W. R.; Johnson, E. C.; Reilly, E. P.

2023-10-17 neuroscience 10.1101/2023.10.16.562590 medRxiv
Top 0.1%
12.2%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWData from a variety of domains are represented as graphs, including social networks, transportation networks, computer networks, and biological networks. A key question spans these domains: are there meaningful repeated subgraphs, or motifs, within the structure of these larger networks? This is a particularly relevant problem when searching for repeated neural circuits in networks of biological neurons, as the field now regularly produces large brain connectivity maps of neurons and synapses, or connectomes. Given acquisition costs, however, these neuron-synapse connectivity maps are mostly one-of-a-kind. With current graph analysis techniques, it is very challenging to discover new "interesting" subgraphs a priori given small sample sizes of host graphs. Another challenge is that for even relatively modest graph sizes, an exhaustive search of all possible subgraphs is computationally intractable. For these reasons, motif discovery in biological graphs remains an unsolved challenge in the field. In this work, we present a motif discovery approach that can derive a list of undirected or directed motifs, with occurrence counts which are statistically significant compared to randomized graphs, from a single graph example. We first address common pitfalls in the current most common approaches when testing for motif statistical significance, and outline a strategy to ameliorate this problem with improved graph randomization techniques. We then propose a progressive-refinement approach for motif discovery, which addresses issues of computational cost. We demonstrate that our sampling correction technique allows for significance testing of target motifs while highlighting misleading conclusions from standard random graph models. Finally, we share our reference implementation, which is available as an open-source Python package, and demonstrate real-world preliminary results on the C. elegans connectome and the ellipsoid body of the Drosophila melanogaster fruit fly connectome.

7
PANDORA: Population Archive of Neuroimaging Data Organized for Rapid Analysis

Abivardi, A.; Webster, M.; McCarthy, P.; Alfaro-Magro, F.; Radosavljevic, L.; Miller, K. L.; Jbabdi, S.; Woolrich, M. W.; Gong, W.; Beckmann, C. F.; Elliott, L. T.; Nichols, T. E.; Smith, S. M.

2026-01-06 radiology and imaging 10.64898/2026.01.05.26343425 medRxiv
Top 0.1%
12.1%
Show abstract

Population-scale neuroimaging allows for novel biological discovery, but voxelwise analyses are computationally paralyzing and noisy, whereas imaging-derived phenotypes discard crucial spatial detail. We introduce PANDORA (Population Archive of Neuroimaging Data Organized for Rapid Analysis), a data-adaptive modelling platform designed to resolve this trade-off. PANDORA has encoded brain MRI data comprising 98 sub-modalities from over 80,000 UK Biobank participants in a highly efficient supervoxel representation. By performing statistical regressions directly within this compressed embedding, PANDORA reduces storage by up to 99% and accelerates computation 10-fold, while acting as a spatial denoiser to enhance statistical power. PANDORA also includes the full-resolution voxelwise ground-truth data, curated imaging confound variables, and a fast analysis tool achieving whole brain, voxelwise population-level regression in seconds to minutes. We showcase PANDORAs ability to reproduce known patterns and reveal new associations including trauma, anxiety/depression, autism polygenic scores, and EPHA3.

8
Inferring Cell Differentiation Dynamics with Unobserved Progenitors

Howard-Snyder, W.; Zhang, R.; Schmidt, H.; Chan, M.; Raphael, B.

2025-12-12 developmental biology 10.64898/2025.12.09.693214 medRxiv
Top 0.1%
11.8%
Show abstract

A cell differentiation map describes the transitions between cell types during a developmental process, and determining this map is a key challenge in developmental biology. Recent single-cell lineage tracing technologies generate lineage trees that describe the history of cell divisions during a developmental process but do not directly measure differentiation events between cell types. Current approaches to infer cell differentiation maps from cell lineage trees make unrealistic assumptions about the developmental process, do not allow for unobserved progenitor cell types, or do not infer cell-type specific rates of growth and transitions. To address these issues, we introduce TROUPE, a likelihood-based framework that infers differentiation and growth dynamics directly from leaf-labeled cell lineage trees while allowing for unobserved progenitor types via biologically motivated potency constraints. We provide an efficient algorithm for computing the maximum likelihood transition and growth rates, as well as a simple model-selection scheme to choose the number of unobserved types. On simulations, TROUPE recovers transition and growth rates better than previous approaches even when there are unobserved types. We apply TROUPE to a model of mammalian development, Trunk-Like Structures (TLS), and show that TROUPE computes reasonable rates of self-renewal and differentiation in both standard and perturbed conditions.

9
Replay of Interictal Sequential Activity Shapes the Epileptic Network Dynamics

Wang, K.; Wang, H.; Yan, Y.; Li, W.; Cai, F.; Zhou, W.; Hong, B.

2024-03-30 neurology 10.1101/2024.03.28.24304879 medRxiv
Top 0.1%
11.8%
Show abstract

Both the imbalance of neuronal excitation and inhibition, and the network disorganization may lead to hyperactivity in epilepsy. However, the insufficiency of seizure data poses the challenge of elucidating the network mechanisms behind the frequent and recurrent abnormal discharges. Our study of two extensive intracranial EEG datasets revealed that the seizure onset zone exhibits recurrent synchronous activation of interictal events. These synchronized discharges formed repetitive sequential patterns, indicative of a stable and intricate network structure within the seizure onset zone (SOZ). We hypothesized that the frequent replay of interictal sequential activity shapes the structure of the epileptic network, which in turn supports the occurrence of these discharges. The Hopfield-Kuramoto oscillator network model was employed to characterize the formation and evolution of the epileptic network, encoding the interictal sequential patterns into the network structure using the Hebbian rule. This model successfully replicated patient-specific interictal sequential activity. Dynamic change of the network connections was further introduced to build an adaptive Kuramoto model to simulate the interictal to ictal transition. The Kuramoto oscillator network with adaptive connections (KONWAC) model we proposed essentially combines two scales of Hebbian plasticity, shaping both the stereotyped propagation and the ictal transition in epileptic networks through the interplay of regularity and uncertainty in interictal discharges.

10
Virtual brain twins for stimulation in epilepsy

Wang, H. E.; Dollomaja, B.; TRIEBKORN, J. P.; Duma, G. M.; WILLIAMSON, A.; Makhalova, J.; LEMARERECHAL, J.-d.; BARTOLOMEI, F.; Jirsa, V.

2024-07-27 neurology 10.1101/2024.07.25.24310396 medRxiv
Top 0.1%
11.7%
Show abstract

Estimating the epileptogenic zone network (EZN) is an important part of the diagnosis of drug-resistant focal epilepsy and plays a pivotal role in treatment and intervention. Virtual brain twins based on personalized whole brain modeling provides a formal method for personalized diagnosis by integrating patient-specific brain topography with structural connectivity from anatomical neuroimaging such as MRI and dynamic activity from functional recordings such as EEG and stereo-EEG (SEEG). Seizures demonstrate rich spatial and temporal features in functional recordings, which can be exploited to estimate the EZN. Stimulation-induced seizures can provide important and complementary information. In our modeling process, we consider invasive SEEG stimulation as the most practical current approach, and temporal interference (TI) stimulation as a potential future approach for non-invasive diagnosis and treatment. This paper offers a virtual brain twin framework for EZN diagnosis based on stimulation-induced seizures. This framework estimates the EZN and validated the results on synthetic data with ground-truth. It provides an important methodological and conceptual basis for a series of ongoing scientific studies and clinical usage, which are specified in this paper. This framework also provides the necessary step to go from invasive to non-invasive diagnosis and treatment of drug-resistant focal epilepsy.

11
Building a small brain with a simple stochastic generative model

Richter, O.; Schneidman, E.

2024-07-04 neuroscience 10.1101/2024.07.01.601562 medRxiv
Top 0.1%
11.7%
Show abstract

The architectures of biological neural networks result from developmental processes shaped by genetically encoded rules, biophysical constraints, stochasticity, and learning. Understanding these processes is crucial for comprehending neural circuits structure and function. The ability to reconstruct neural circuits, and even entire nervous systems, at the neuron and synapse level, facilitates the study of the design principles of neural systems and their developmental plan. Here, we investigate the developing connectome of C. elegans using statistical generative models based on simple biological features: neuronal cell type, neuron birth time, cell body distance, reciprocity, and synaptic pruning. Our models accurately predict synapse existence, degree profiles of individual neurons, and statistics of small network motifs. Importantly, these models require a surprisingly small number of neuronal cell types, which we infer and characterize. We further show that to replicate the experimentally-observed developmental path, multiple developmental epochs are necessary. Validation of our models predictions of the synaptic connections using multiple reconstructions of adult worms suggests that our model identified the fundamental "backbone" of the connectivity graph. The accuracy of the generative statistical models we use here offers a general framework for studying how connectomes develop and the underlying principles of their design.

12
Constructing genotype and phenotype network helps reveal disease heritability and phenome-wide association studies

Cao, X.; Zhu, L.; Liang, X.; Zhang, S.; Sha, Q.

2023-11-20 genetic and genomic medicine 10.1101/2023.11.14.23297400 medRxiv
Top 0.1%
11.7%
Show abstract

Analyses of a bipartite Genotype and Phenotype Network (GPN), linking the genetic variants and phenotypes based on statistical associations, provide an integrative approach to elucidate the complexities of genetic relationships across diseases and identify pleiotropic loci. In this study, we first assess contributions to constructing a well-defined GPN with a clear representation of genetic associations by comparing the network properties with a random network, including connectivity, centrality, and community structure. Next, we construct network topology annotations of genetic variants that quantify the possibility of pleiotropy and apply stratified linkage disequilibrium (LD) score regression to 12 highly genetically correlated phenotypes to identify enriched annotations. The constructed network topology annotations are informative for disease heritability after conditioning on a broad set of functional annotations from the baseline-LD model. Finally, we extend our discussion to include an application of bipartite GPN in phenome-wide association studies (PheWAS). The community detection method can be used to obtain a priori grouping of phenotypes detected from GPN based on the shared genetic architecture, then jointly test the association between multiple phenotypes in each network module and one genetic variant to discover the cross-phenotype associations and pleiotropy. Significance thresholds for PheWAS are adjusted for multiple testing by applying the false discovery rate (FDR) control approach. Extensive simulation studies and analyses of 633 electronic health record (EHR)-derived phenotypes in the UK Biobank GWAS summary dataset reveal that most multiple phenotype association tests based on GPN can well-control FDR and identify more significant genetic variants compared with the tests based on UK Biobank categories.

13
Multiscale Hyperbolic Embedding for Cell Hierarchies in Large-Scale Bioinformatics Data

Yao, M.; Praturu, A.; Sharpee, T.

2025-10-01 biophysics 10.1101/2025.09.29.679407 medRxiv
Top 0.1%
11.6%
Show abstract

The increasing size of datasets poses challenges for their visualization and interpretation, highlighting the need for scalable and effective analysis methods. Hyperbolic embedding have shown strong potential in capturing complex hierarchical structures across diverse systems. However, existing hyperbolic embedding methods typically operate with fixed curvature and have difficulties scaling to large datasets. To address these limitations, we propose MuH-MDS, a novel multiscale algorithm for hyperbolic multidimensional scaling that uses "adiabatic" approximation from physics to optimize local positions while keeping cluster centroid fixed. MuH-MDS improves computing time by 103 compared to previous methods and is able to handle large datasets comprising over 80, 000 samples. We validate the method on a number of datasets, including a large-scale C. elegans embryogenesis scRNA-seq dataset with over 80,000 samples. Here, MuH-MDS uncovers intrinsic hierarchical structure, and achieves improved pseudotime inference and lineage analysis compared to UMAP and other methods. Unlike UMAP and t-SNE, which emphasize local structure at the expense of global coherence and metric accuracy, MuH-MDS preserves global hierarchy in a metrically faithful manner, maintaining key relationships across scales.

14
Simulation-Based Inference for Whole-Brain Network Modeling of Epilepsy using Deep Neural Density Estimators

Hashemi, M.; Vattikonda, A. N.; Jha, J.; Sip, V.; Woodman, M. M.; Bartolomei, F.; Jirsa, V.

2022-06-03 neurology 10.1101/2022.06.02.22275860 medRxiv
Top 0.1%
10.9%
Show abstract

Whole-brain network modeling of epilepsy is a data-driven approach that combines personalized anatomical information with dynamical models of abnormal brain activity to generate spatio-temporal seizure patterns as observed in brain imaging signals. Such a parametric simulator is equipped with a stochastic generative process, which itself provides the basis for inference and prediction of the local and global brain dynamics affected by disorders. However, the calculation of likelihood function at whole-brain scale is often intractable. Thus, likelihood-free inference algorithms are required to efficiently estimate the parameters pertaining to the hypothetical areas in the brain, ideally including the uncertainty. In this detailed study, we present simulation-based inference for the virtual epileptic patient (SBI-VEP) model, which only requires forward simulations, enabling us to amortize posterior inference on parameters from low-dimensional data features representing whole-brain epileptic patterns. We use state-of-the-art deep learning algorithms for conditional density estimation to retrieve the statistical relationships between parameters and observations through a sequence of invertible transformations. This approach enables us to readily predict seizure dynamics from new input data. We show that the SBI-VEP is able to accurately estimate the posterior distribution of parameters linked to the extent of the epileptogenic and propagation zones in the brain from the sparse observations of intracranial EEG signals. The presented Bayesian methodology can deal with non-linear latent dynamics and parameter degeneracy, paving the way for reliable prediction of neurological disorders from neuroimaging modalities, which can be crucial for planning intervention strategies.

15
Combined statistical-mechanistic modeling links ionchannel genes to physiology of cortical neuron types

Bernaerts, Y.; Deistler, M.; Goncalves, P. J.; Beck, J.; Stimberg, M.; Scala, F.; Tolias, A. S.; Macke, J. H.; Kobak, D.; Berens, P.

2023-03-02 neuroscience 10.1101/2023.03.02.530774 medRxiv
Top 0.1%
10.8%
Show abstract

Neural cell types have classically been characterized by their anatomy and electrophysiology. More recently, single-cell transcriptomics has enabled an increasingly fine genetically defined taxonomy of cortical cell types, but the link between the gene expression of individual cell types and their physiological and anatomical properties remains poorly understood. Here, we develop a hybrid modeling approach to bridge this gap. Our approach combines statistical and mechanistic models to predict cells electrophysiological activity from their gene expression pattern. To this end, we fit biophysical Hodgkin-Huxley-based models for a wide variety of cortical cell types using simulation-based inference, while overcoming the challenge posed by the mismatch between the mathematical model and the data. Using multimodal Patch-seq data, we link the estimated model parameters to gene expression using an interpretable sparse linear regression model. Our approach recovers specific ion channel gene expressions as predictive of biophysical model parameters including ion channel densities, directly implicating their mechanistic role in determining neural firing.

16
All-atom protein design via SE(3) flow matching with ProteinZen

Li, A. J.; Kortemme, T.

2025-10-18 bioengineering 10.1101/2025.10.18.683228 medRxiv
Top 0.1%
10.5%
Show abstract

The advent of generative models of protein structure has greatly accelerated our ability to perform de novo protein design, especially concerning design at coarser physical scales such as backbone generation and protein binder design. How-ever, the design of precise placements at atomic scales remains a challenge for existing design methods. One avenue towards higher fidelity atomic-scale design is via generative models with full atomic resolution, but is complicated by the intricacies of simultaneously designing both discrete protein sequence and continuous atomic positionings. In this work we propose a framework to capture this interplay by decomposing residues into collections of oriented rigid bodies, allowing us to apply SE(3) flow-matching for all-atom protein structure generation. Our method, ProteinZen1, generates designs with high sequence-structure consistency while retaining competitive diversity and novelty on both unconditional and conditional generation tasks. We demonstrate competitive performance for unconditional monomer design and state-of-the-art performance on various forms of motif scaffolding, including full-atom motif scaffolding and motif scaffolding without specifying motif segment spacing or relative sequence order.

17
Personalized Feature Statistics: Individual-Level Variant Inference under Genetic Ancestry Continuum

Wang, J. F.; Yu, R.; Edelson, J.; Park, J.; Le Guen, Y.; Liu, X.; Belloy, M.; Ionita-Laza, I.; Greicius, M.; Tang, H.; He, Z.

2026-04-29 neurology 10.64898/2026.04.28.26351879 medRxiv
Top 0.1%
10.1%
Show abstract

Genome-wide association studies (GWAS) have successfully identified numerous genetic variants associated with complex diseases. However, the extent to which the effects of these variants vary across populations of diverse ancestries remains poorly understood. Furthermore, in these contexts genetic ancestry is treated as a categorical variable, thereby oversimplifying its continuous nature and the more nuanced ways in which it can influence genetic effects on disease. Here, we propose personalized feature statistics (PFstatistics), a statistical framework that quantifies the importance of genetic variants to a phenotype based on each individuals ancestry background, and profiles heterogeneous genetic effects across the genetic ancestry continuum. We demonstrate the utility of this framework through both simulations and real data analysis using sequencing data from ancestrally diverse cohorts in the Alzheimers Disease Sequencing Project (ADSP). We show that Alzheimers Disease (AD) risk variants span a spectrum from ancestry-homogeneous to ancestry-dependent effects, and that PFstatistics characterizes this spectrum at individual resolution across the ancestry continuum. PFstatistics also provides individual-level variant selection with FDR controlled at a target level, yielding distinct selection sets that vary across individuals according to their ancestry background. While demonstrated in the context of genetic ancestry, the proposed method is broadly applicable to other heterogeneity features such as environmental factors, offering a robust tool for understanding complex genetic contributions across diverse populations.

18
Integrative multi-omics QTL colocalization maps regulatory architecture in aging human brain

Cao, X.; Sun, H.; Feng, R.; Mazumder, R.; Najar, C. F. B. A.; Li, Y. I.; De Jager, P. L.; Bennett, D. A.; The Alzheimer's Disease Functional Genomics Consortium, ; Dey, K. K.; Wang, G.

2025-04-20 neurology 10.1101/2025.04.17.25326042 medRxiv
Top 0.1%
9.9%
Show abstract

Multi-trait QTL (xQTL) colocalization has shown great promises in identifying causal variants with shared genetic etiology across multiple molecular modalities, contexts, and complex diseases. However, the lack of scalable and efficient methods to integrate large-scale multi-omics data limits deeper insights into xQTL regulation. Here, we propose ColocBoost, a multi-task learning colocalization method that can scale to hundreds of traits, while accounting for multiple causal variants within a genomic region of interest. ColocBoost employs a specialized gradient boosting framework that can adaptively couple colocalized traits while performing causal variant selection, thereby enhancing the detection of weaker shared signals compared to existing pairwise and multi-trait colocalization methods. We applied ColocBoost genome-wide to 17 gene-level single-nucleus and bulk xQTL data from the aging brain cortex of ROSMAP individuals (average N = 595), encompassing 6 cell types, 3 brain regions and 3 molecular modalities (expression, splicing, and protein abundance). Across molecular xQTLs, ColocBoost identified 16,503 distinct colocalization events, exhibiting 10.7({+/-} 0.74)-fold enrichment for heritability across 57 complex diseases/traits and showing strong concordance with element-gene pairs validated by CRISPR screening assays. When colocalized against Alzheimers disease (AD) GWAS, ColocBoost identified up to 2.5-fold more distinct colocalized loci, explaining twice the AD disease heritability compared to fine-mapping without xQTL integration. This improvement is largely attributable to ColocBoosts enhanced sensitivity in detecting gene-distal colocalizations, as supported by strong concordance with known enhancer-gene links, highlighting its ability to identify biologically plausible AD susceptibility loci with underlying regulatory mechanisms. Notably, several genes including BLNK and CTSH showed sub-threshold associations in GWAS, but were identified through multi-omics colocalizations which provide new functional support for their involvement in AD pathogenesis.

19
DarkQ: Continuous genomic monitoring using message queues

Viehweger, A.; Brandt, C.; Hölzer, M.

2020-11-13 genomics 10.1101/2020.11.12.379560 medRxiv
Top 0.1%
9.9%
Show abstract

MotivationNewly sequenced genomes are often not noticed by potential stakeholders because submission to public databases is delayed, and search options are limited. However, the discovery of genomes can be vital: In pathogen outbreaks, fast updates are essential to coordinate containment efforts and prevent spread. ResultsHere we introduce DarkQ, a message queue that allows for instant sharing and discovery of genomes. AvailabilityDarkQ is released under the BSD-2 license at github.com/phiweger/darkq.

20
TRIBAL: Tree Inference of B cell Clonal Lineages

Weber, L. L.; Reiman, D.; Roddur, M. S.; Qi, Y. L.; El-Kebir, M.; Khan, A. A.

2023-11-27 immunology 10.1101/2023.11.27.568874 medRxiv
Top 0.1%
9.8%
Show abstract

B cells are a critical component of the adaptive immune system, responsible for producing antibodies that help protect the body from infections and foreign substances. Single cell RNA-sequencing (scRNA-seq) has allowed for both profiling of B cell receptor (BCR) sequences and gene expression. However, understanding the adaptive and evolutionary mechanisms of B cells in response to specific stimuli remains a significant challenge in the field of immunology. We introduce a new method, TRIBAL, which aims to infer the evolutionary history of clonally related B cells from scRNA-seq data. The key insight of TRIBAL is that inclusion of isotype data into the B cell lineage inference problem is valuable for reducing phylogenetic uncertainty that arises when only considering the receptor sequences. Consequently, the TRIBAL inferred B cell lineage trees jointly capture the somatic mutations introduced to the B cell receptor during affinity maturation and isotype transitions during class switch recombination. In addition, TRIBAL infers isotype transition probabilities that are valuable for gaining insight into the dynamics of class switching. Via in silico experiments, we demonstrate that TRIBAL infers isotype transition probabilities with the ability to distinguish between direct versus sequential switching in a B cell population. This results in more accurate B cell lineage trees and corresponding ancestral sequence and class switch reconstruction compared to competing methods. Using real-world scRNA-seq datasets, we show that TRIBAL recapitulates expected biological trends in a model affinity maturation system. Furthermore, the B cell lineage trees inferred by TRIBAL were equally plausible for the BCR sequences as those inferred by competing methods but yielded lower entropic partitions for the isotypes of the sequenced B cell. Thus, our method holds the potential to further advance our understanding of vaccine responses, disease progression, and the identification of therapeutic antibodies. AvailabilityTRIBAL is available at https://github.com/elkebir-group/tribal